Papers with content moderation strategies
MOSAIC: Modeling Social AI for Content Dissemination and Regulation in Multi-Agent Simulations (2025.emnlp-main)
Copied to clipboard
| Challenge: | generative language agents predict user behaviors such as liking, sharing, and flagging content. |
| Approach: | They propose a framework where generative language agents predict user behaviors such as liking, sharing, and flagging content. |
| Outcome: | The proposed framework analyzes content moderation strategies and user engagement dynamics at scale and demonstrates that agents’ articulated reasoning for their social interactions aligns with their collective engagement patterns. |
Conspiracy Theories and Where to Find Them on TikTok (2025.acl-long)
Copied to clipboard
| Challenge: | Existing studies on TikTok's potential to promote and amplify harmful content have not been conducted. |
| Approach: | They analyze a longitudinal dataset of 1.5M videos shared in the U.S. over three years and evaluate the effects of TikTok’s Creativity Program for monetization. |
| Outcome: | The proposed model achieves high precision in detecting harmful content, but its overall performance is comparable to fine-tuned traditional models such as RoBERTa. |
Unmasking the Imposters: How Censorship and Domain Adaptation Affect the Detection of Machine-Generated Tweets (2025.coling-main)
Copied to clipboard
| Challenge: | generative AI has been used to generate fluent and convincing text on social media platforms . a new study examines the generative capabilities of four popular large language models . |
| Approach: | They propose a methodology to examine the generative capabilities of four prominent LLMs on Twitter using a dataset from Llama 3, Mistral, Qwen2 and GPT4o. |
| Outcome: | The proposed method examines the generative capabilities of four prominent LLMs on Twitter. |
FigSIM: A Dataset for Fine-grained Suicide Severity and Figurative Language in Suicide Memes (2026.findings-acl)
Copied to clipboard
| Challenge: | Suicide memes are increasingly common on social media, yet remain poorly understood and potentially harmful. |
| Approach: | They propose a dataset designed for fine-grained analysis of suicide memes and benchmark 16 models for figurative language, suicide severity, and content detection. |
| Outcome: | The proposed model outperforms existing models on figurative language, suicide severity, and suicide-related content detection tasks. |